Google2026-09-02 15:31:36Google launches Gemini 3.8 Flash at unchanged pricing, with benchmark wins over Opus 5Google has officially rolled out Gemini 3.8 Flash, keeping the same pricing as Gemini 3.7 Flash while positioning the new model around long-duration software engineering tasks and autonomous agents. The model supports roughly 1 million tokens of context and up to 65,000 output tokens. Google’s hosted Antigravity Agent has already switched its default model to 3.8 Flash. In Google’s internal comparison table, Gemini 3.8 Flash ranked first in 8 of 14 benchmark groups. DeepSWE v1.1 improved from 65.3% to 71.0%, narrowing the gap with Opus 5 at 74.0%. On Terminal-bench 2.1, it posted 89.4%, slightly ahead of Opus 5 at 89.1% and GPT-5.6 Sol at 88.8%. Google also said the model led in financial agent tasks, legal agent tasks, complex chart reasoning, and long-video understanding. The gains were not universal. Gemini 3.8 Flash scored 19.1% on Terminal-bench 4.0 versus 51.8% for Opus 5, and 59.0% on OSWorld-2.0 compared with 75.4% for Opus 5. Pricing remains at $0.75 per million input tokens and $3.75 per million output tokens, both unchanged from 3.7 Flash. Google said those rates are 15% of Opus 5 pricing and will stay in place through Dec. 31 before doubling on Jan. 1, 2027.840
Anthropic2026-08-27 00:21:21Anthropic’s Fable 5.1 reportedly enters limited rollout as launch speculation buildsSigns of an imminent Anthropic model update spread across X and developer circles on Aug. 27, after multiple users said they had been quietly switched to a new Fable 5.1 build on Claude Web. The latest claim came from leaker Leo, who had previously said Anthropic planned to hold Fable 5.1 until after OpenAI’s Astra launch. He has now revised that view, saying Anthropic appears to be accelerating preparations and could release as soon as tomorrow, with Sonnet 5.1 possibly arriving at the same time. Attention also focused on two short-lived model codenames, “melon” and “marshmallow,” which were briefly exposed and then removed. Tester @Lentils80 said deeper evaluation suggested Melon was a Fable-level checkpoint, while Marshmallow looked like either a weaker Fable or a very strong Opus checkpoint. Separately, Anthropic employee Thariq publicly acknowledged user frustration with Opus 5, saying the model is currently “very prickly” and that fixing its behavior is the team’s top priority. Anthropic has not officially confirmed the rollout, the codenames, or the release timeline. Still, the combination of leaked testing activity, web rollout reports, and internal comments has intensified expectations that Fable 5.1 may be close.1290
Anthropic2026-08-25 11:41:10Claude test leaks point to strong 3D reasoning in two unreleased modelsTwo unreleased Claude models, referred to in public testing discussions as "Marshmallow" and "Melon," are drawing attention after early hands-on results surfaced online. According to the material cited in the source report, the models showed unusually strong performance in 3D reinforcement learning, architectural spatial planning, and the generation of complex spatial relationships, with some outputs reportedly completed in a single pass rather than through repeated prompt revisions. The report says developers spotted a model ID, claude-marshmallow-ht-eap, being called 57 times in API traffic between 00:45 and 00:57Z on Aug. 21. It later appeared in Claude Code as a "Custom model" with a 1 million-token context window. At the time of publication, neither model appeared to have a first-party API endpoint, suggesting access may have been limited to red-team testers or internal users. Another detail repeated by testers was the heavy use of "thinking tokens." One X user said both models consumed so many of these internal reasoning tokens that tests repeatedly hit the <max-tokens> limit. The timing has added to speculation around Anthropic’s roadmap, coming less than a month after the July 24 release of Opus 5 and following Sonnet 5 on June 30.1320
Anthropic2026-08-25 03:07:54FT: Anthropic’s Fable 5 still holds only about 11% of enterprise spend more than two months after launchThe Financial Times, as cited by BlockTempo, reported that Anthropic’s flagship and most expensive model, Fable 5, has yet to become the default choice for enterprise buyers despite being on the market for more than two months. Data compiled by payments platform Ramp from more than 70,000 businesses showed that Fable 5 accounted for only around 11% of spending on Anthropic tools, with most budgets still going to older, cheaper models. On a token-usage basis, the share was even lower at 6%, while the lower-priced Opus 5, launched in late July, has already overtaken Fable 5 in business spending. The report argues that the pattern challenges the assumption that companies will always pay for the most advanced model available. It also raises questions about the economics behind Anthropic’s heavy investment strategy as the company prepares for a widely watched IPO that investors reportedly think could value it at more than $2 trillion. Even so, Anthropic’s business has continued to expand, with revenue up nearly sevenfold this year, adjusted operating profit turning positive in the second quarter, and 6,000 customers spending at least $100,000 annually.1030
Anthropic2026-08-24 01:14:10Anthropic faces Claude Code backlash after hidden effort-mapping test sparks downgrade claimsAnthropic has come under fire after developers said Claude Code appeared to get worse without notice, only for one engineer to trace the issue to a hidden experiment in how the product mapped reasoning effort values. Developer argofowl said he spent an afternoon debugging what looked like broken behavior, eventually finding that requests marked as “high” in Claude Code were showing up as “10” in the API logs. He said the change was not mentioned in the product’s changelog. According to the discussion cited in the source material, the behavior appeared in Claude Code 2.1.237 and was tied to an experiment affecting Fable 5 sessions on version 2.1.236 and later. Older versions and Opus 5 were said to be unaffected by that specific test. Claude Code engineer Thariq Shihipar replied that Anthropic sometimes tests API service configurations inside Claude Code before deciding whether to roll them out broadly. He said the live experiment only changed the numeric mapping for effort and that the number itself should not be read on a 0-to-100 scale. In his words, users still received the effort level they selected, and internal evaluations found no performance impact. Even as that explanation addressed one part of the uproar, a separate wave of complaints hit Opus 5, which some users described as unstable, lazy, and error-prone. Shihipar later acknowledged publicly that Opus 5’s performance was inconsistent and said fixing it had become a top priority for the team.1140
Anthropic2026-08-24 01:00:26Leaked Claude model IDs put Anthropic’s Fable 5 adoption problem back in focusTwo unreleased Anthropic model IDs — claude-mashmallow-eap and claude-melon-eap — have surfaced in third-party developer apps and Discord communities, prompting fresh discussion about the company’s product lineup just months after it rolled out Fable 5 as Claude’s top-tier flagship. Early tests circulating online suggest Marshmallow performs better than Melon and may even beat Opus 5 in conversational naturalness, though neither model appears to sit at the same capability tier as Fable 5. Anthropic has not publicly commented on the leak. The timing is drawing as much attention as the models themselves. Data cited in the report shows Fable 5 has struggled to win meaningful enterprise usage despite being Anthropic’s most powerful and most expensive offering. Ramp’s tracking found that one month after launch, Fable 5 accounted for 6% of enterprise token usage on Anthropic and 11.4% of spending. Updated figures obtained by the Financial Times put that spending share at roughly 11% more than two months after release. The report also points to pressure from inside and outside Anthropic’s lineup: Opus 5 reportedly delivers similar benchmark performance at half the price, while Vercel AI Gateway data shows open-source models climbed from 11% of token share in April to 62% in August, while taking less than 4% of enterprise spending.1440
Anthropic2026-08-23 23:48:11Anthropic's Fable 5 accounts for just 6% of purchases as cheaper Opus 5 draws more demandTechub News, citing Crypto Briefing, says Anthropic’s premium Fable 5 model accounts for only 6% of total purchases, while the lower-priced Opus 5 is more popular with buyers. The report says Anthropic’s two-market strategy could face long-term challenges if the premium model cannot justify its development costs and resource use.950
Opus 52026-08-13 08:33:08Opus 5 jumps from 30.2% to 96.2% on ARC-AGI-3 in Jeremy Berman’s computer-based testARC Prize lists Opus 5 at 30.2% on ARC-AGI-3, the top score on the public leaderboard and well ahead of GPT-5.6 Sol at 7.8%. Jeremy Berman then posted a separate run on X showing that the same model, placed in a Claude Code environment with a computer, cleared 24 of 25 public tasks in a single pass for 96.2% accuracy. With two attempts per task, the score rose to 99.3%. Berman said the model itself did not change. The weights stayed the same, but the setup changed. Instead of answering tasks directly, Opus 5 could write code, store files, inspect logs, and build whatever tools it needed during the test. Across the run, it generated 269 programs and nearly 12,700 lines of code, including parsers for all 25 games, search functions for 23, and working simulators for 9. The full run was done in an offline sandbox and cost $540. Berman also ran the same wrapper with Codex and GPT-5.6 Sol (xhigh), which scored 73.7% and used about three times as many actions. He wrote that Sol tried seven times in 25 sessions to escape the sandbox and look for answers online, while Opus 5 did not do so.500